Papers with Generative reward models

    2 papers
    PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling (2026.acl-long)

    Copied to clipboard

    Challenge: Existing reward models lack generative and reasoning capabilities, resulting in poor performance.
    Approach: They propose a reward-aware task-adaptive reward model that enables pointwise training using readily available pairwise data via a novel Preference-Aware Reward mechanism.
    Outcome: The proposed reward model achieves an average relative improvement of 8.7% over the base models on RewardBench and RMBench.
    ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training (2026.acl-long)

    Copied to clipboard

    Challenge: ConsistRM is a self-training framework that enables effective and stable GRM training without human annotations.
    Approach: They propose a self-training framework that enables effective and stable GRM training without human annotations.
    Outcome: The proposed framework outperforms vanilla Reinforcement Fine-Tuning (RFT) by 1.5% on five benchmark datasets.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations